Journal of Molecular Biology
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Journal of Molecular Biology's content profile, based on 232 papers previously published here. The average preprint has a 0.13% match score for this journal, so anything above that is already an above-average fit.
Rigoli, M.; Faccioli, P.; Biasini, E.
Show abstract
Prion diseases are fatal neurodegenerative disorders driven by the conversion of the cellular prion protein (PrP) into a misfolded, pathogenic conformer. Beyond serving as a substrate for prion propagation, PrP is also thought to mediate neurotoxic signaling. Within this framework, the central region of PrP has emerged as a critical regulatory element. Notably, deletion of residues 105-125 ({Delta}CR) leads to spontaneous neurodegeneration in vivo and induces abnormal ionic currents in cultured cells and primary neurons, indicating that this region is essential for controlling the toxicity of the N-terminal domain. Current models propose that the N-terminus functions as a toxic effector whose activity is modulated by the C-terminal domain. This intramolecular interplay is likely central to the physiological role of PrP, and its disruption may contribute to neurodegeneration. Here, we investigated how deletion of the central region affects the structure and dynamics of full-length PrP. We generated membrane-bound models of full-length, diglycosylated wild-type (WT) PrP and the neurotoxic {Delta}CR mutant, and compared their conformational dynamics using molecular dynamics simulations. The two proteins exhibited markedly distinct behaviours. WT PrP adopted a more compact conformational ensemble of the N-terminal domain, consistent with stabilizing interactions between the flexible N-terminus and the globular C-terminal domain. In contrast, the {Delta}CR variant displayed more extended conformations and a substantial redistribution of intramolecular contacts, including the loss of specific interactions between the disordered N-terminal tail and the globular domain. This altered structural organization was accompanied by an increased propensity of the N-terminal domain to approach the membrane surface in the mutant. Our results provide a molecular model in which the central region engages intramolecular interaction networks that ultimately help regulate N-terminal residence at the plasma membrane, offering mechanistic insight into how CR deletion shifts the conformational ensemble toward membrane-associated states that may be associated with neurotoxic activity.
Rawat, P.; Ramakrishnan, P.; Cardente, N.; Kumar, S.; Greiff, V.; Gromiha, M. M.
Show abstract
Protein aggregation is central to amyloid-related disorders and remains a major developability challenge for protein therapeutics. Over the past two decades, significant advances have been made to predict aggregation-prone regions (APRs) and estimate aggregation propensity in proteins and peptides. In contrast, the prediction of aggregation kinetics has received relatively less attention due to the limited availability and heterogeneity of experimental data. Consequently, aggregation propensities from APR prediction algorithms were widely accepted as a means to predict relative changes in the aggregation kinetics of proteins and mutants. Previous studies have demonstrated, using large-scale datasets, that aggregation propensity shows a weak or inconsistent correlation with aggregation kinetics. In the present study, we have integrated complementary state-of-the-art mechanistic and kinetic prediction tools for protein aggregation into a unified, user-friendly web framework entitled "Amylo-Pipe". Amylo-Pipe also implements practical features that are especially useful for protein engineering, such as gatekeeper-residue mutational scanning to support the design of aggregation-resistant variants. By consolidating multiple prediction tasks in a single interface, Amylo-Pipe enables a more comprehensive assessment of aggregation behavior than APR-only workflows. The web server is freely accessible at: https://web.iitm.ac.in/bioinfo2/amylopipe/.
Mead, E. H.; Batz, K. C.; Shih, K.-H.; Fleming, I. R.; Tesdahl, C. D.; Lizardos, L.; Armendariz, J. R.; Hannan, J. P.; Hickey, A. M.; Leyk, A.; Erbse, A. H.; Falke, J. J.
Show abstract
The three conventional isoforms of the Ras G-protein (H-, K-, N-Ras) function as molecular on-off switches that regulate a wide array of signaling pathways, including the Ras-PI3K-PIP3-PDK1-AKT pathway that is central to innate immunity and normal cell growth, and is dysregulated in many disease states. Activation of the pathway by Ras requires adequate Ras-PI3K binding affinity. Here we focus on the interface of known structure in the H-Ras:PI3K{gamma} co-complex essential to multiple pathways including directed pseudopod growth in leukocyte chemotaxis. At this interface 10 H-Ras residues, all 100% conserved between the H-, K- and N-Ras isomers, contact the Ras binding domain of PI3K{gamma} (PI3K{gamma}RBD). To investigate the degree to which the native H-Ras:PI3K{gamma}RBD interface is optimized by evolution for maximal binding affinity, 8 interfacial Ras mutations selected from the COSMIC database and the literature were introduced at the contact positions. All 8 Ras mutations were observed to alter the H-Ras:PI3K{gamma}RBD binding affinity, with 4 mutations yielding significant affinity increases and 4 yielding significant affinity decreases. These findings indicate that the native H-Ras:PI3K{gamma}RBD interface provides intermediate, rather than maximal, binding affinity. Such intermediate affinity is consistent with the substantial binding plasticity of the conserved H-, N-, K-Ras effector docking surface, which has evolved to bind a diverse array of effectors. Furthermore, the findings provide evidence that COSMIC-linked mutations at the H-Ras:PI3K{gamma}RBD interface frequently generate affinity increases as well as decreases, with potential implications for molecular mechanisms of disease and for tool development in cell biology.
Olivieri, F.;Konstantinova, A.;Ribnikar, N.;Bizjak, N.;Žnidar, ?.;Abel, K.;Rajh, E.;Ljubetič, A.
Show abstract
Over the past decade, protein design has evolved from a specialized discipline into a broadly accessible approach for engineering and interrogating biological systems. Despite these advances, protein design continues to be a technically challenging task, often requiring knowledge of programming to be able to use and combine the different software packages. To address this challenge, we have developed Prosculpt, an easy-to-use protein design pipeline. Prosculpt integrates RFdiffusion for backbone generation, ProteinMPNN for sequence design and multiple structure-prediction platforms (AF2, AF3, Colabfold, Boltz2). Candidate designs are evaluated using customizable Rosetta-based scoring protocols. Each project is specified through a single configuration file, enabling users with minimal computational expertise to perform sophisticated protein design tasks without writing code, while also allowing advanced users to access the full capabilities of the underlying programs. Prosculpt supports a wide range of applications, including design of symmetric homo-oligomers, design of binders, motif scaffolding, partial diffusion and fixed-backbone sequence redesign. By combining these capabilities within a single, user-friendly platform, Prosculpt provides a practical entry point to modern protein design for both novice and expert users.
Mureddu, L. G.; Brooksbank, E. J.; Vuister, G. W.; Muskett, F. W.
Show abstract
Nuclear Magnetic Resonance (NMR) relaxation experiments provide a powerful residue-resolved access to biomolecular dynamics across a wide range of timescales. Unfortunately, the quantitative analysis of the relaxation data remains distributed across specialised and often disconnected tools. Here, we present CcpNmr AnalysisDynamics, the latest addition to the CcpNmr Analysis program suite, providing an integrated framework for relaxation analysis, exchange-aware interpretation and dynamical modelling. The platform unifies relaxation-rate extraction, diagnostic validation, model-based analysis and structural visualisation within reproducible workflows, while supporting future extension through a robust application programming interface and plugin architecture. We introduce ModelAnalysis (ModA), a new analysis engine based on the Lipari-Szabo formalism that incorporates robust optimisation, uncertainty estimation and model-selection strategies designed for heterogeneous relaxation datasets. The framework also supports exchange-focused analysis and integration with specialised external modelling tools, allowing relaxation anomalies to be followed from initial detection to more detailed interpretation. The applicability and reliability of AnalysisDynamics are demonstrated through systematic re-analysis and validation of curated relaxation datasets from the Biological Magnetic Resonance Data Bank. These analyses enable assessment of data consistency, dynamic parameters and model reliability across magnetic fields, providing a reproducible route from NMR relaxation measurements to structure-linked interpretation of biomolecular dynamics.
Nde, J.; Panapitiya, G.; Cheung, M. S.; Maupin, C. M.; Sardiu, M. E.
Show abstract
The INO80 chromatin remodeling complex plays a central role in DNA repair, transcription, and replication. Yet, a comprehensive understanding of its structural organization remains incomplete due to the dynamic nature of several of its subunits and the sharing of several subunits with related remodeling complexes. Here, we report a computational model of the three-dimensional structure of the S. cerevisiae INO80 complex using an integrative approach that combines experimental crosslinking mass spectrometry, molecular docking, and molecular dynamics simulations. Our results reveal the spatial and dynamical organization of key modules--ARP8, ARP5, NHP10, and RVB1/2--within the intact complex. The resulting structural model agrees with crosslinking constraints, highlighting the architecture of the previously uncharacterized NHP10 module. This module, including the C-terminal region of the Ino80 scaffolding protein, has remained elusive due to its intrinsic flexibility and lack of high-resolution structural data. To facilitate this integrative modeling workflow and make it broadly accessible, we presented INTEGRATOR: (INTEGRAtive TempOral and stRuctural Analysis of protein modules), a versatile workflow package designed as a tool to elucidate the structure and dynamics of large, flexible macromolecular assemblies using well-established softwares. Our findings demonstrate the power of integrative modeling in resolving the role of the highly disordered NPH10 module in recruiting other dynamic modules into INO80 large protein assemblies and offer a generalizable framework for determining the architecture of similarly complex and heterogeneous molecular machines. This work carries broad implications for understanding the structural basis of chromatin regulation in microbial organisms and the implications for the dysregulation in diseases such as cancer.
Nimkar, S.; Nguyen, T.; Karandur, D.; Subramanian, S.; O'Donnell, M. E.; Kuriyan, J.
Show abstract
DNA polymerase clamp loaders are AAA+ ATPases that load sliding clamps on DNA for high- speed replication. Using a platform for high-throughput mutagenesis of replication proteins in T4 bacteriophage, we carried out saturation mutagenesis of the AAA+ ATPase module of the T4 clamp loader bearing a mutation, Gln 118{lozenge}Asn (Q118N), that reduces fitness. We identified residues for which different mutations improve the fitness of the Q118N variant but are neutral in the wild-type background. These conditionally neutral "rescue hotspots" overlap with those identified earlier in another defective variant (D110C). These rescue hotspots localize to regions where the sequence is not optimal for the structure, as determined by energetic frustration analysis. We designed new sequences for three of these regions, using the protein-design algorithm ProteinMPNN. In two helical regions, several designed sequences increased the fitness of both wild-type and mutant proteins, likely due to enhanced stability. An inter-domain hinge in AAA+ module changes conformation during activation, and designs for the hinge lead to loss of fitness in the wild-type background. However, when using the active conformation as the template, designs for the hinge increase the fitness of defective variants. In contrast designs templated on the inactive conformation led to loss of fitness, suggesting that a proper conformational balance is crucial. Thus, adaptive capacity in the clamp loader resides in a network of conditionally neutral sites that enable functional tuning through shifts in stability and conformational equilibria.
Benavides, T. L.; Ramelot, T. A.; Montelione, G. T.
Show abstract
While allosteric protein function has been appreciated for decades, the ubiquity of conformational shifts, particularly those distant from the interaction interface, has not been broadly characterized. For example, ligand binding frequently triggers allosteric effects far from the interaction interface, yet the prevalence of these conformational shifts underpinning protein function remain poorly documented. We systematically assessed the generality of allosteric effects as monitored by NMR Chemical Shift Perturbations (CSPs) distant from the interaction interface. In a set of 139 protein-protein complexes, a striking 74% of all significant CSPs are non-local to the binding site. Notably, more than 35% of significant CSPs outside the binding site occur in residues for which the shortest receptor-ligand interatomic distance is more than 10 [A]. Every protein analyzed exhibits a significant fraction (> 8%) of CSPs distant from the binding site. This analysis across a large number of protein structures demonstrates and documents that structural plasticity is a ubiquitous and fundamental property of proteins. Significance StatementStudies of protein dynamics have had a profound impact on biology. Ruth Nussinov famously postulated that multiple protein conformations preexist in dynamic equilibrium, with interconversions that mediate function. While conformational flexibility has been characterized in many specific case studies, the extent to which structural plasticity can be considered a fundamental and ubiquitous property of proteins remains poorly documented. We address a central question: how common is protein structural plasticity? To do so, we compiled a database of protein-protein and protein-peptide complexes with NMR chemical shift data for both bound (holo) and unbound (apo) states. These data reveal the widespread prevalence of long-range structural perturbations induced by ligand binding, demonstrating that structural plasticity is a pervasive and fundamental property of proteins.
Liu, Z. H.; Zhang, O.; De Castro, S.; Sun, K.; Ghafouri, H.; Attafi, O. A.; Fawzi, N. L.; Tosatto, S. C. E.; Monzon, A. M.; Moses, A. M.; Head-Gordon, T.; Forman-Kay, J. D.
Show abstract
More than two thirds of proteins in the human proteome are predicted to contain intrinsically disordered regions (IDRs), which lack stable folded structure. IDRs are critical for biological regulation and organization, as targets for post-translational modifications, and as mediators of biomolecular condensates. To address the pressing need for better structural models enabling functional insight, we developed AlphaFlex to model fully atomistic conformer ensembles for proteins predicted to have IDRs, modeled in the context of AlphaFold folded domains and an implicit bilayer for transmembrane proteins. The AlphaFlex resource provides conformational ensembles of human proteins from the AlphaFold database with identified IDRs in the Protein Ensemble Database that is mirrored in UniProt. This transformative resource of AlphaFlex ensembles provides physically and biologically relevant full-length models for IDR proteins, including scaffold proteins, those with IDR:folded-domain interactions, regulatory and condensate proteins requiring exposed binding elements, conditionally folding IDRs, and transmembrane proteins containing IDRs.
Adkins, B. J.; Sidlowski, P. F. W.; Jennings, C. E.; Morrison, E. A.
Show abstract
Nuclear organization is dynamic and originates from the fundamental subunit of chromatin, the nucleosome. Post-translational modification of nucleosomal histones, particularly within intrinsically disordered histone tail regions, provides a dynamic regulatory mechanism of accessibility for chromatin-templated processes. While the epigenomic impacts of lysine acetylation and serine phosphorylation in the histone H3 tail are well-known, how these charge-altering post-translational modifications (PTMs) alter nucleosomal tail conformational dynamics remains incompletely characterized. Given that the functional implications of these PTMs are, at least in part, a consequence of modified nucleosome conformation, systematically cataloging the impact of histone PTMs on nucleosome dynamics provides crucial insight into both baseline cellular activity and epigenetic dysregulation that occurs in disease. Previously, our lab demonstrated that arginine citrullination mimetics lead to regional increases in H3 tail dynamics within nucleosome core particles. Here, we performed nuclear magnetic resonance spin relaxation experiments to investigate the effects of lysine acetylation and serine phosphorylation on H3 tail picosecond-nanosecond (ps-ns) dynamics. Using lysine-to-glutamine and serine-to-glutamate mutations as acetyllysine and phosphoserine mimetics, respectively, we found that these PTMs increase ps-ns conformational dynamics regionally around the PTM site, with a position-dependent effect. Additionally, we show that the type of PTM influences the extent of these increases: in general, the effect of mimetics trends in the order of phosphorylation [≤] acetylation < citrullination, suggesting a tunable method for altering histone tail dynamics. Taken together, these results illustrate the role of nucleosome conformational dynamics in conveying the effects of epigenomic PTMs, elucidating a mechanism of the histone language.
Kim, A.-R.; Perrimon, N.
Show abstract
As protein structure prediction tools become widely adopted across biology, there is a growing need for accessible methods to assess and visualize predicted protein-protein interactions (PPIs). Here we present LIVIA (Local Interaction Visualization and Analysis), a browser-based tool that computes local PPI confidence metrics across multiple prediction platforms, identifies predicted interface residues, embeds an interactive Mol* 3D viewer, and generates visualization scripts for ChimeraX and PyMOL. The tool automatically detects prediction formats; all parsing and computation occur locally on the users machine. LIVIA is freely available at https://flyark.github.io/LIVIA.
Gharaie Amirabadi, D.; Jackson, C.; Kim, D. S.; Sprang, M.; Amani, K.
Show abstract
Protein engineering often relies on separate models for related developability properties, limiting efficiency and transfer across tasks. We present Prot2Prop, a multitask framework based on a frozen ProstT5 encoder with shared and task-specific adapters for joint prediction of six protein properties: material production, solubility, temperature stability, aggregation propensity, expression yield, and folding stability. Across held-out test data, Prot2Prop achieved strong performance on both classification and regression tasks, including AUROC values ranging from 0.86 to 0.98 for classification endpoints and Spearman correlations ranging from 0.73 to 0.86 for regression endpoints. The model achieved particularly strong performance for temperature stability (AUROC = 0.98) and aggregation propensity (Spearman = 0.86). Post-hoc calibration further improved regression accuracy, reducing folding stability MAE from 0.67 to 0.48. These results demonstrate that parameter-efficient multitask adaptation of protein language models can provide accurate and unified prediction of diverse protein developability properties. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=132 SRC="FIGDIR/small/735009v1_ufig1.gif" ALT="Figure 1"> View larger version (51K): org.highwire.dtl.DTLVardef@d0eea6org.highwire.dtl.DTLVardef@e3f482org.highwire.dtl.DTLVardef@1c98656org.highwire.dtl.DTLVardef@192a93b_HPS_FORMAT_FIGEXP M_FIG C_FIG
Chen, A.; Siddiqui, J.; Taucar, W.; Tiralongo, L.; Tkachenko, M.; Xu, A.; Bawa, S.; Guo, S.; Pinska, O.; Rim, J.; Shi, J.; Wang, M.; Zhao, E.
Show abstract
Inverse-folding models can rapidly generate protein sequences compatible with a supplied backbone, but unconstrained redesign is poorly suited to enzyme and genome-editor-associated domains, where catalytic, substrate-proximal, and conserved structural regions must remain protected. In this paper, we present EditorForge, a modular constraint-and-audit suite for editor-domain protein redesign that wraps fixed-backbone inverse folding with explicit design masks, fixed-position enforcement, active-site-proximity auditing, active-site-shielded regeneration, and downstream structural quality control. Using full-length Moloney murine leukemia virus reverse transcriptase structure 4MH8 (MMLV RT 4MH8) as a demonstration target, EditorForge first restricted redesign to a bounded 25-position envelope while fixing 428 residues. An initial audit detected active-site-proximal failure modes despite fixed-position integrity. Later, the Active Site Shield module then removed five unsafe design positions, replaced them with lower-contact alternatives, and regenerated candidates under stricter constraints. Post Shield Audit evaluated 24 regenerated candidates, all of which satisfied the hard sequence/mask and active-site-shield constraints. For the eight candidates that were selected or returned for structure-prediction/refolding quality control, Enhanced RefoldQC found that all 8 evaluated predicted structures passed the computational structure-QC screen. That said, the selected 8 candidates passed the computational structure-QC screen, with global C RMSD values of 1.2061-1.5555 {degrees}A, active-site C RMSD values of 0.4098-1.8397 {degrees}A, mutation-neighborhood C RMSD values of 1.3155-1.6848 {degrees}A, and average pLDDT-like confidence values of 94.87-95.11. In short, EditorForge provides a reproducible triage layer that converts general inverse-folding output into constrained and editor-specific candidate sets for downstream structural and biological review on top of existing structural prediction tools.
Ploskon, E.; Baskaran, K.; Tejero, R.; Schwieters, C. D.; Bardiaux, B.; Guentert, P.; Fogh, R. H.; Gutmanas, A.; Brooksbank, E. J.; Yokochi, M.; Wishart, D. S.; Wedell, J. R.; Vranken, W. F.; Thompson, D.; Thompson, G.; Smith, B. O.; Rehman, S.; Ramelot, T. A.; Ragan, T. J.; Perez, A.; Perera, B. L.; Peisach, E.; Nilges, M.; Mureddu, L. G.; Mondal, A.; Lubicka, E. A.; Liwo, A.; Kurisu, G.; Kobayashi, N.; Klukowski, P.; Johnston, B. A.; Huang, Y. J.; Hoch, J. C.; Higman, V. A.; Herrmann, T.; Hayward, M. W.; Garnet, J. A.; Case, D. A.; Burley, S. K.; Adams, P. D.; Montelione, G. T.; Vuister, G.
Show abstract
The NMR Exchange Format (NEF) is a community-driven standard for representing NMR experimental data in a consistent, interoperable, and machine-readable form. Built on the STAR syntax, NEF provides a structured framework for storing and exchanging chemical shifts, peak lists, various types of structural restraints, and related metadata, thus allowing for data exchange across software platforms. By enabling direct, lossless transfer of information, NEF simplifies multi-software workflows, improves reproducibility, and supports FAIR (Findable, Accessible, Interoperable, Reusable) data principles. We describe the NEF specification, its current implementation across commonly used NMR software packages, and its application in areas including biomolecular structure determination, metabolomics, and ligand screening. Testing demonstrates that NEF can be used to exchange complete datasets between programs without loss of information or functionality. We also outline recent developments and future directions, such as inclusion of NMR relaxation data and support for non-standard residue topologies. NEFs growing adoption highlights its potential as a unifying standard for NMR data, enabling more efficient, transparent and collaborative research.
Smyth, S.; Liu, Z. H.; Tsangaris, T.; Head-Gordon, T.; Forman-Kay, J. D.; Gradinaru, C. C.
Show abstract
Eukaryotic cap-dependent translation initiation is regulated by binding of the predominantly folded eukaryotic initiation factor 4E (eIF4E) to the intrinsically disordered eIF4E binding proteins (4E-BPs). Here, we report full-length atomistic conformational ensembles generated by IDPConformerGenerator and optimized by X-EISDv2 workflow for both apo 4E-BP2, the neuronal 4E-BP, and 4E-BP2 in complex with eIF4E, using data from single-molecule fluorescence and nuclear magnetic resonance (NMR), together with select coordinates from a 4E-BP1:eIF4E crystal structure. Structural sampling within dynamic complexes is often under-appreciated, with NMR and crystal structure data for 4E-BP:eIF4E suggesting different degrees of structural heterogeneity. Our ensemble models validated by solution spectroscopy data enable comparison of free 4E-BP2 and its complex with eIF4E. This shows a delocalization of contacts around canonical regions, which supports previous findings of unidirectional conditional occupancy of the binding sites. Two new contact regions emerged: one between the disordered N-termini of eIF4E and 4E-BP2, which may play an allosteric role in tuning the binding affinity, and the other between the C-terminus of 4E-BP2 and an extended region of eIF4E, which is consistent with the extended, dynamic binding interface that we reported previously. These results support a model of translation regulation in which the dynamic 4E-BP2:eIF4E complex facilitates accessibility of regulatory sites of 4E-BP2 when bound.
Stankus, M.; Anderson, M.
Show abstract
Human glutathione synthetase (hGS) is a negatively cooperative ATP-grasp enzyme that catalyzes the final step in the biosynthesis of glutathione, a tripeptide antioxidant critical for life. hGS functions as an obligate homodimer with one active site per subunit; the two active sites are separated by [~]40 Angstroms. How ligand binding in one subunit reshapes the distant partner active site has remained a central unresolved question in understanding hGS regulation. This study provides the first atomistic model of ligand-dependent inter-subunit communication underlying negative cooperativity in hGS. Using atomistic simulations and dynamical network analysis, this study reveals how reactant- and product-bound states remodel the empty partner active site, redistribute inter-subunit interactions, and organize long-range communication between the two active sites. The product-bound/partner-empty state displayed a larger and less hydrated empty active site, demonstrating that ligand identity in one subunit alters both the geometry and solvent environment of the opposite site. Changes in ligand-dependent interactions are distributed across the dimer interface, with prominent contributions from the 42-46 interface region, the 11-30 region, and the 212-236 helical/interface region. Suboptimal path analysis shows product- and reactant-bound states share a communication scaffold, with 64.1% of transmission residues common to both pathways, 30.8% product-specific, and 5.1% reactant-specific. Together, the present results establish a detailed structural framework for hGS negative cooperativity in which ligand binding remodels a distributed allosteric network linking substrate-binding loops, the dimer interface, and the partner active site. More broadly, this work demonstrates how atomistic simulations can resolve long-range active-site coupling in multimeric enzymes and provides a foundation for experimental tests of allosteric transmission in hGS.
Bergsma, T.; Kolbe Musskopf, M.; Feito, A.; Gallardo, P.; Rebeaud, M. E.; Kuiper, E. F.; Hernandez Espejo, N.; Tejedor, A. R.; Feenstra, J.; Fernando, S. M. Y.; Steen, A.; Vlijm, R.; Espinosa, J. R.; Kampinga, H.; Veenhoff, L.
Show abstract
Molecular chaperones are known for their role in preventing protein aggregation and assisting proteins in reaching their structurally functional state. DNAJB6, a J-domain protein that partners with Hsp70s and nucleotide exchange factors, is very potent in preventing amyloid formation of proteins with large intrinsically disordered regions (IDRs), including several disease-associated proteins. Complementary to this, we recently demonstrated a role for DNAJB6 in surveilling FG-Nucleoporins (FG-Nups) phase transitions and highlighted its role in nuclear pore complex assembly. We expand on this by showing that this activity of phase state surveillance is directed to several FG-Nups and shared with the closely related DNAJB2 and DNAJB8. We demonstrate that the surveillance mechanism of DNAJB6 is encoded in an unusually highly conserved IDR that promotes the formation of stable, gel-like assemblies of the chaperone itself. These assemblies likely provide a stable environment that can outcompete stable homotypic FG-Nup interactions and instead favors multivalent heterotypic chaperone:FG-Nup interactions. The evolutionary conservation of the DNAJB6-IDR, mutant analyses from both experimental in vitro and in cell data, and multiscale molecular dynamics simulations suggest that the sequence space for encoding stable gel-like assemblies is narrow and optimized to avoid self-aggregation while providing potent anti-amyloidogenic capacity.
Castillo, S.; Gu, C.; Jouhten, P.; Peddinti, G.; Ollila, S. O. H.
Show abstract
Accurate enzyme function annotation remains a major bottleneck in genome analysis despite the rapid expansion of available protein sequence and structure data. Most existing methods rely on sequence similarity or machine-learning representations, which often perform poorly for proteins with low sequence identity or convergent evolutionary histories. Because enzymatic activity is determined by the three-dimensional arrangement of catalytic and binding-site residues, structure-based approaches offer a mechanistically grounded alternative. However, their broader application has been constrained by the limited size and coverage of curated active-site reference databases. To address this challenge, we developed ActSeekN, a structural-motif-based functional annotation pipeline that combines the ActSeek active-site search algorithm with a newly constructed large-scale reference database derived from AlphaFold-predicted structures, UniProt annotations, and curated catalytic residue information. This framework enables rapid and scalable identification of conserved catalytic motifs across structurally related proteins, allowing function to be transferred on the basis of local three-dimensional catalytic geometry rather than global sequence similarity. In this way, ActSeekN overcomes a central limitation of previous structure-based methods by expanding the searchable space of catalytic motifs while retaining mechanistic interpretability. Benchmarking against state-of-the-art machine-learning approaches demonstrates competitive or superior performance. Applications to yeast, human, and Trichoderma reesei proteomes refine existing annotations, complete partial EC assignments, and identify previously unrecognized enzymatic functions, highlighting ActSeekN as a powerful tool for genome annotation and biotechnology. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=99 SRC="FIGDIR/small/720574v1_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@f5da16org.highwire.dtl.DTLVardef@c0faa4org.highwire.dtl.DTLVardef@18765fforg.highwire.dtl.DTLVardef@3960b9_HPS_FORMAT_FIGEXP M_FIG C_FIG
Bellaiche, A.; Choudhary, P.; Nair, S.; Harrus, D.; Yu, C. W.-H.; Tanweer, S. A.; Evans, G. L.; Lo, S. W.; Martin, M.; Fleming, J. R.; Velankar, S.
Show abstract
Structure Integration with Function, Taxonomy and Sequences (SIFTS) provides residue-level mappings between UniProt Knowledgebase sequences and Protein Data Bank structures and has historically been generated through internal Protein Data Bank in Europe (PDBe) pipelines. Here, PDBe-SIFTS is presented as a fully open-source, locally deployable implementation of this mapping framework. The pipeline combines fast, scalable sequence search using MMseqs2, an improved bounded scoring scheme for ranking candidate mappings, and residue-level mapping refinement based on backbone connectivity. PDBe-SIFTS is distributed as a Python package with command-line tools for 1) building a sequence search database, 2) identifying the best sequence-structure match, 3) one-to-one mapping at the residue level, and 4) generating SIFTS annotations in PDBx/mmCIF format. Benchmarking on the complete Protein Data Bank archive showed that MMseqs2 reduced archive-scale UniProtKB searches from hours with BLASTP to minutes, approximately 22-36 times faster, while curated mappings were recovered at top rank in 93.1% of cases. The remaining discrepancies mainly involved biologically ambiguous cases such as highly conserved proteins, chimeric constructs, or closely related orthologs. These results show that PDBe-SIFTS enables fast mapping, improving structural coherence in residue-level alignments while delivering the most up-to-date and accurate mappings, comparable to expert curation. Tool: https://github.com/PDBeurope/SIFTS Quick start notebook with example: https://github.com/PDBeurope/SIFTS/tree/master/notebooks Broader audience statementMatching protein sequences to their three-dimensional structures, and mapping annotations across both, is essential for understanding protein function, interactions, and molecular mechanisms. This integrated view enables richer interpretation of biological data and underpins advances in drug discovery, disease research, and protein engineering. PDBe-SIFTS provides an open and functional framework for structure-sequence mapping, allowing researchers and databases to run, inspect, and extend these mappings locally, while benefiting from faster searches, transparent scoring, and structurally informed residue-level alignments. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=110 SRC="FIGDIR/small/721839v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@5e6ea6org.highwire.dtl.DTLVardef@1b2754dorg.highwire.dtl.DTLVardef@1334f9forg.highwire.dtl.DTLVardef@1b083a1_HPS_FORMAT_FIGEXP M_FIG C_FIG
Villalona, P.; Pulahinge, T.; Yu, T.; Wenning, J.; Frisbie, C. J.; Magafas, J.; Okafor, C. D.
Show abstract
The nuclear receptor superfamily is comprised of ligand-regulated transcription factors that contain an intrinsically disordered domain at the amino-terminal end, known as the N-terminal domain (NTD). While this poorly conserved domain is known to possess ligand-independent activation function (AF-1), few NTD functions are conserved between nuclear receptors (NRs). Identified roles in other receptors include androgen receptor (AR), estrogen receptor (ER) and mineralocorticoid receptor (MR). Here, we aim to define the function of the NTD of the farnesoid X receptor (FXR), a crucial regulator of lipid and bile acid metabolism. We show that the NTD engages in interdomain contact with other FXR domains. We also observe that the NTD interacts directly with coregulator proteins. Using mutagenesis, mammalian two-hybrid assays and molecular dynamics simulations, we identify and validate a novel SXXLF motif in the NTD which mediates interactions with both coregulators and the ligand binding domain. Mutation of the motif induces large changes in conformational and allosteric coupling in FXR. Our study identifies a new nuclear receptor-interacting motif that modulates the transcriptional activity of FXR. Graphical AbstractFXR-NTD regulates transcriptional activity through interdomain communication with the LBD and is also involved in co-activator recruitment. The SENLF motif is the first defined functional element within the FXR-NTD and mediates both NTD-LBD interaction and selective co-activator engagements to drive NTD-mediated transcriptional activity. O_FIG O_LINKSMALLFIG WIDTH=135 HEIGHT=200 SRC="FIGDIR/small/724725v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@5a37aorg.highwire.dtl.DTLVardef@2fa9e1org.highwire.dtl.DTLVardef@13a19daorg.highwire.dtl.DTLVardef@1775ed2_HPS_FORMAT_FIGEXP M_FIG C_FIG